Skip to content

feat(parquet): add wide-schema writer overhead benchmark - #9723

Merged
alamb merged 1 commit into
apache:mainfrom
HippoBaro:wide_schema_writer_bench
Apr 15, 2026
Merged

feat(parquet): add wide-schema writer overhead benchmark#9723
alamb merged 1 commit into
apache:mainfrom
HippoBaro:wide_schema_writer_bench

Conversation

@HippoBaro

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

Existing writer benchmarks use narrow schemas (5–10 columns) and primarily measure data encoding throughput. They don't capture per-column structural overhead that dominates at high column cardinality (thousands to hundreds of thousands of columns), such as allocation, and metadata assembly.

What changes are included in this PR?

This commit adds benchmarks to fill that gap by writing a single-row batch through ArrowWriter with 1k/5k/10k flat Float32 columns and per-column WriterProperties entries, isolating the cost of the writer infrastructure itself.

Baseline results (Apple M1 Max):

  writer_overhead/1000_cols/per_column_props      3.72 ms
  writer_overhead/5000_cols/per_column_props     54.96 ms
  writer_overhead/10000_cols/per_column_props   220.73 ms

Are these changes tested?

N/A

Are there any user-facing changes?

N/A

Existing writer benchmarks use narrow schemas (5–10 columns) and
primarily measure data encoding throughput. They don't capture
per-column structural overhead that dominates at high column cardinality
(thousands to hundreds of thousands of columns), such as allocation, and
metadata assembly.

This commit adds benchmarks to fill that gap by writing a single-row
batch through `ArrowWriter` with 1k/5k/10k flat `Float32` columns and
per-column `WriterProperties` entries, isolating the cost of the writer
infrastructure itself.

Baseline results (Apple M1 Max):

  writer_overhead/1000_cols/per_column_props      3.72 ms
  writer_overhead/5000_cols/per_column_props     54.96 ms
  writer_overhead/10000_cols/per_column_props   220.73 ms

Signed-off-by: Hippolyte Barraud <hippolyte.barraud@datadoghq.com>
@github-actions github-actions Bot added the parquet Changes to the parquet crate label Apr 15, 2026

@etseidl etseidl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good. The metadata bench will do wide tables (10k columns), but only measures decoding the footer. Nice to have something similar on the write side.

@alamb
alamb merged commit 06c3bd0 into apache:main Apr 15, 2026
17 checks passed
@alamb

alamb commented Apr 15, 2026

Copy link
Copy Markdown
Contributor

Thank you @HippoBaro and @etseidl for the review

alamb added a commit that referenced this pull request Apr 16, 2026
# Which issue does this PR close?

- Depends on #9723
- Contributes to #9722

# Rationale for this change

`WriterProperties::offset_index_disabled()` checked whether any column
in the `column_properties` HashMap has page-level statistics enabled,
scanning the entire map on every call. This method is called from
`GenericColumnWriter::new` — once per column per row group. With N
columns each having per-column properties, this resulted in quadratic
HashMap iterations during row group construction.

# What changes are included in this PR?

Move the scan into `WriterPropertiesBuilder::build()` so it runs once at
construction time.

Benchmark results (vs baseline in #9723):

```
  writer_overhead/1000_cols/per_column_props     2.44 ms  (was   3.25 ms, −25%)
  writer_overhead/5000_cols/per_column_props    13.28 ms  (was  47.45 ms, −72%)
  writer_overhead/10000_cols/per_column_props   27.97 ms  (was 197.97 ms, −86%)
```

Scaling now linear.

# Are these changes tested?

All tests passing.

# Are there any user-facing changes?

None.

Signed-off-by: Hippolyte Barraud <hippolyte.barraud@datadoghq.com>
Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
Rich-T-kid pushed a commit to Rich-T-kid/arrow-rs that referenced this pull request Jun 2, 2026
# Which issue does this PR close?

- Contributes to apache#9722

# Rationale for this change

Existing writer benchmarks use narrow schemas (5–10 columns) and
primarily measure data encoding throughput. They don't capture
per-column structural overhead that dominates at high column cardinality
(thousands to hundreds of thousands of columns), such as allocation, and
metadata assembly.

# What changes are included in this PR?

This commit adds benchmarks to fill that gap by writing a single-row
batch through `ArrowWriter` with 1k/5k/10k flat `Float32` columns and
per-column `WriterProperties` entries, isolating the cost of the writer
infrastructure itself.

Baseline results (Apple M1 Max):

```
  writer_overhead/1000_cols/per_column_props      3.72 ms
  writer_overhead/5000_cols/per_column_props     54.96 ms
  writer_overhead/10000_cols/per_column_props   220.73 ms
```

# Are these changes tested?

N/A

# Are there any user-facing changes?

N/A

Signed-off-by: Hippolyte Barraud <hippolyte.barraud@datadoghq.com>
Rich-T-kid pushed a commit to Rich-T-kid/arrow-rs that referenced this pull request Jun 2, 2026
…he#9724)

# Which issue does this PR close?

- Depends on apache#9723
- Contributes to apache#9722

# Rationale for this change

`WriterProperties::offset_index_disabled()` checked whether any column
in the `column_properties` HashMap has page-level statistics enabled,
scanning the entire map on every call. This method is called from
`GenericColumnWriter::new` — once per column per row group. With N
columns each having per-column properties, this resulted in quadratic
HashMap iterations during row group construction.

# What changes are included in this PR?

Move the scan into `WriterPropertiesBuilder::build()` so it runs once at
construction time.

Benchmark results (vs baseline in apache#9723):

```
  writer_overhead/1000_cols/per_column_props     2.44 ms  (was   3.25 ms, −25%)
  writer_overhead/5000_cols/per_column_props    13.28 ms  (was  47.45 ms, −72%)
  writer_overhead/10000_cols/per_column_props   27.97 ms  (was 197.97 ms, −86%)
```

Scaling now linear.

# Are these changes tested?

All tests passing.

# Are there any user-facing changes?

None.

Signed-off-by: Hippolyte Barraud <hippolyte.barraud@datadoghq.com>
Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

parquet Changes to the parquet crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants